<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Document processing</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Document_processing"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Document_processing rootpage-Document_processing skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Document processing</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p><b>Document processing</b> is a field of research and a set of <a href="Production_process" class="mw-redirect" title="Production process">production processes</a> aimed at making an analog <a href="Document" title="Document">document</a> digital. Document processing does not simply aim to photograph or <a href="Image_scanning" class="mw-redirect" title="Image scanning">scan</a> a document to obtain a <a href="Digital_image" title="Digital image">digital image</a>, but also to make it digitally intelligible. This includes extracting the structure of the document or the <a href="Document_layout_analysis" title="Document layout analysis">layout</a> and then the content, which can take the form of text or images. The process can involve traditional <a href="Computer_vision" title="Computer vision">computer vision</a> algorithms, convolutional neural networks or manual labor. The problems addressed are related to <a href="Semantic_segmentation" class="mw-redirect" title="Semantic segmentation">semantic segmentation</a>, <a href="Object_detection" title="Object detection">object detection</a>, <a href="Optical_character_recognition" title="Optical character recognition">optical character recognition (OCR)</a>, <a href="Handwritten_text_recognition" class="mw-redirect" title="Handwritten text recognition">handwritten text recognition (HTR)</a> and, more broadly, <a href="Transcription_(linguistics)" title="Transcription (linguistics)">transcription</a>, whether <a href="Automation" title="Automation">automatic</a> or not.<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup> The term can also include the phase of digitizing the document using a scanner and the phase of interpreting the document, for example using <a href="Natural_language_processing" title="Natural language processing">natural language processing</a> (NLP) or <a href="Image_classification" class="mw-redirect" title="Image classification">image classification</a> technologies. It is applied in many industrial and scientific fields for the optimization of administrative processes, mail processing and the digitization of analog <a href="Archiving" class="mw-redirect" title="Archiving">archives</a> and historical documents.
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Background">Background</h2></div>
<p>Document processing was initially as is still to some extent a kind of production line work dealing with the treatment of <a href="Document" title="Document">documents</a>, such as letters and parcels, in an aim of sorting, extracting or massively extracting data. This work could be performed in-house or through <a href="Business_process_outsourcing" title="Business process outsourcing">business process outsourcing</a>.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> Document processing can indeed involve some kind of externalized manual labor, such as <a href="Amazon_Mechanical_Turk" title="Amazon Mechanical Turk">mechanical Turk</a>.
</p><p>As an example of manual document processing, as relatively recent as 2007,<sup id="cite_ref-VisaDox_4-0" class="reference"><a href="#cite_note-VisaDox-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup> document processing for "millions of visa and citizenship applications" was about use of "approximately 1,000 contract workers" working to "manage mail room and <a href="Data_entry_clerk" title="Data entry clerk">data entry</a>."
</p><p>While document processing involved data entry via keyboard well before use of a <a href="Computer_mouse" title="Computer mouse">computer mouse</a> or a <a href="Image_scanner" title="Image scanner">computer scanner</a>, a 1990 article in <i><a href="The_New_York_Times" title="The New York Times">The New York Times</a></i> regarding what it called the "<a href="Paperless_office" title="Paperless office">paperless office</a>" stated that "document processing begins with the scanner".<sup id="cite_ref-Paper.NYT_5-0" class="reference"><a href="#cite_note-Paper.NYT-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> In this context, a former <a href="Xerox" title="Xerox">Xerox</a> vice-president, <a href="Paul_Strassmann" title="Paul Strassmann">Paul Strassmann</a>, expressed a critical opinion, saying that computers add rather than reduce the volume of paper in an office.<sup id="cite_ref-Paper.NYT_5-1" class="reference"><a href="#cite_note-Paper.NYT-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup> It was said that the engineering and maintenance documents for an airplane weigh "more than the airplane itself".
</p>
<div class="mw-heading mw-heading2"><h2 id="Automatic_document_processing">Automatic document processing</h2></div>
<p>As the <i><a href="State_of_the_art" title="State of the art">state of the art</a></i> advanced, document processing transitioned to handling "document components ... as database entities."<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</p><p>A technology called automatic document processing or sometimes intelligent document processing (IDP) emerged as a specific form of <a href="Process_Automation" class="mw-redirect" title="Process Automation">Intelligent Process Automation</a> (IPA), combining <a href="Artificial_intelligence" title="Artificial intelligence">artificial intelligence</a> such as <a href="Machine_Learning" class="mw-redirect" title="Machine Learning">Machine Learning</a> (ML), <a href="Natural_Language_Processing" class="mw-redirect" title="Natural Language Processing">Natural Language Processing</a> (NLP) or <a href="Intelligent_Character_Recognition" class="mw-redirect" title="Intelligent Character Recognition">Intelligent Character Recognition</a> (ICE) to extract data from several types documents.<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup> Advancements in automatic document processing, also called Intelligent Document Processing, improve the ability to process <a href="Unstructured_data" title="Unstructured data">unstructured data</a> with fewer exceptions and greater speeds. <sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Applications">Applications</h3></div>
<p>Automatic document processing applies to a whole range of documents, whether structured or not. For instance, in the world of business and finance, technologies may be used to process paper-based invoices, forms, purchase orders, contracts, and currency bills.<sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup> Financial institutions use intelligent document processing to process high volumes of forms such as regulatory forms or loan documents. ID uses AI to extract and classify data from documents, replacing manual data entry.<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup>
</p><p>In medicine, document processing methods have been developed to facilitate patient follow-up and streamline administrative procedures, in particular by digitizing medical or laboratory analysis reports. The goal is also to standardize medical databases.<sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup> Algorithms are also directly used to assist physicians in medical diagnosis, e.g. by analyzing <a href="Magnetic_resonance_imaging" title="Magnetic resonance imaging">magnetic resonance images</a>,<sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup> or <a href="Microscope" title="Microscope">microscopic</a> images.<sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup>
</p><p>Document processing is also widely used in the <a href="Humanities" title="Humanities">humanities</a> and <a href="Digital_humanities" title="Digital humanities">digital humanities</a>, in order to extract historical <a href="Big_data" title="Big data">big data</a> from archives or heritage collections. Specific approaches were developed for various sources, including textual documents, such as newspaper archives,<sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup> but also images,<sup id="cite_ref-cini_archive_digitization_17-0" class="reference"><a href="#cite_note-cini_archive_digitization-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> or maps.<sup id="cite_ref-18" class="reference"><a href="#cite_note-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Technologies">Technologies</h3></div>
<p>If, from the 1980s onward, traditional computer vision algorithms were widely used to solve document processing problems,<sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup> these have been gradually replaced by neural network technologies in the 2010s.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup> However, traditional computer vision technologies are still used, sometimes in conjunction with neural networks, in some sectors.
</p><p>Many technologies support the development of document processing, in particular <a href="Optical_character_recognition" title="Optical character recognition">optical character recognition</a> (OCR), and <a href="Handwritten_text_recognition" class="mw-redirect" title="Handwritten text recognition">handwritten text recognition</a> (HTR), which allow the text to be transcribed automatically. Text segments as such are identified using instance or <a href="Object_detection" title="Object detection">object detection</a> algorithms, which can sometimes also be used to detect the structure of the document. The resolution of the latter problem sometimes also uses <a href="Semantic_segmentation" class="mw-redirect" title="Semantic segmentation">semantic segmentation</a> algorithms.
</p><p>These technologies often form the core of document processing. However, other algorithms may intervene before or after these processes. Indeed, document <a href="Digitization" title="Digitization">digitization</a> technologies are also involved, whether in the form of classical or three-dimensional scanning.<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup> The digitization of 3D documents can in particular resort to derivatives of <a href="Photogrammetry" title="Photogrammetry">photogrammetry</a>. Sometimes, specific 2D scanners must also be developed to adapt to the size of the documents or for reasons of scanning ergonomics.<sup id="cite_ref-cini_archive_digitization_17-1" class="reference"><a href="#cite_note-cini_archive_digitization-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup> The document processing also depends on the digital encoding of the documents in a suitable <a href="File_format" title="File format">file format</a>. Furthermore, the processing of heterogeneous databases can rely on <a href="Image_classification" class="mw-redirect" title="Image classification">image classification</a> technologies.
</p><p>At the other end of the chain are various image completion, extrapolation or data cleanup algorithms. For textual documents, the interpretation can use <a href="Natural_language_processing" title="Natural language processing">natural language processing</a> (NLP) technologies.
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="Document_automation" title="Document automation">Document automation</a></li>
<li><a href="Document_modelling" class="mw-redirect" title="Document modelling">Document modelling</a></li>
<li><a href="Data_Processing" class="mw-redirect" title="Data Processing">Data Processing</a></li>
<li><a href="Document_Imaging" class="mw-redirect" title="Document Imaging">Document Imaging</a></li>
<li><a href="Duplex_scanning" title="Duplex scanning">Duplex scanning</a></li>
<li><a href="Text_mining" title="Text mining">Text mining</a></li>
<li><a href="Workflow" title="Workflow">Workflow</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFLen_AspreyMichael_Middleton2003" class="citation book cs1">Len Asprey; Michael Middleton (2003). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=gYOpFlMXcs0C&q=%22document+processing%22+ocr&pg=PA368"><i>Integrative Document & Content Management: Strategies for Exploiting Enterprise Knowledge</i></a>. Idea Group Inc (IGI). <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>9781591400554</bdi>.</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite id="CITEREFVinod_V._Sople2009" class="citation book cs1">Vinod V. Sople (2009-05-25). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=g4dxNB05dgoC&q=document+processing+bpo&pg=PA47"><i>Business Process Outsourcing: A Supply Chain of Expertises</i></a>. PHI Learning Pvt. Ltd. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-8120338159</bdi>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite id="CITEREFMark_Kobayashi-Hillary2005" class="citation book cs1">Mark Kobayashi-Hillary (2005-12-05). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=zdxbEwgfQzQC&q=%22document+processing%22+bpo&pg=PA167"><i>Outsourcing to India: The Offshore Advantage</i></a>. Springer Science & Business Media. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>9783540247944</bdi>.</cite></span>
</li>
<li id="cite_note-VisaDox-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-VisaDox_4-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFJulia_Preston2007" class="citation news cs1">Julia Preston (December 2, 2007). <a rel="nofollow" class="external text" href="https://www.nytimes.com/2007/12/02/us/02immig.html">"Immigration Contractor Trims Wages"</a>. <i><a href="The_New_York_Times" title="The New York Times">The New York Times</a></i>.</cite></span>
</li>
<li id="cite_note-Paper.NYT-5"><span class="mw-cite-backlink">^ <a href="#cite_ref-Paper.NYT_5-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Paper.NYT_5-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFLawrence_M._Fisher1990" class="citation news cs1">Lawrence M. Fisher (July 7, 1990). <a rel="nofollow" class="external text" href="https://www.nytimes.com/1990/07/07/business/paper-once-written-off-keeps-a-place-in-the-office.html">"Paper, Once Written Off, Keeps a Place in the Office"</a>. <i><a href="The_New_York_Times" title="The New York Times">The New York Times</a></i>.</cite></span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text"><cite id="CITEREFAl_YoungDayle_WoolsteinJay_Johnson1996" class="citation magazine cs1">Al Young; Dayle Woolstein; Jay Johnson (February 1996). "Unknown Title". <i>Object Magazine</i>. p. 51.</cite></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="http://www.di.uniba.it/~ndm/pubs/esposito05icdar.pdf">"Intelligent Document processing"</a> <span class="cs1-format">(PDF)</span>. <i>Department of Computer Science – University of Bari</i>. 2005-04-07<span class="reference-accessdate">. Retrieved <span class="nowrap">2018-09-08</span></span>.</cite></span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text"><cite id="CITEREFFloriana_Esposito,_Stefano_Ferilli,_Teresa_M._A._Basile,_Nicola_Di_Mauro2005" class="citation book cs1"><a href="Floriana_Esposito" title="Floriana Esposito">Floriana Esposito</a>, Stefano Ferilli, Teresa M. A. Basile, Nicola Di Mauro (2005-04-01). <a rel="nofollow" class="external text" href="https://www.computer.org/csdl/proceedings-article/icdar/2005/24201100/12OmNqIQS59"><i>"Intelligent Document Processing" in Proceedings. Eighth International Conference on Document Analysis and Recognition, Seoul, South Korea, 2005 pp. 1100-1104. doi: 10.1109/ICDAR.2005.144</i></a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FICDAR.2005.144">10.1109/ICDAR.2005.144</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:17302169">17302169</a>.</cite><span class="cs1-maint citation-comment"><code class="cs1-code">{{cite book}}</code>: CS1 maint: multiple names: authors list (link)</span></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.keymarkinc.com/intelligent-document-processing-idp/">"Intelligent Document Processing (IDP)"</a>. <i>keymarkinc.com</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2024-07-12</span></span>.</cite></span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-10">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1041539562">
/* start https://en.wikipedia.org/ */
.mw-parser-output .citation{word-wrap:break-word}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}
/* end https://en.wikipedia.org/ */
</style><span class="citation patent" id="CITEREFJohn_E._JonesWilliam_J._JonesFrank_M._Csultis2011"><a rel="nofollow" class="external text" href="https://patents.google.com/patent/US7873576B2/en">US active US7873576B2</a>, John E. Jones; William J. Jones & Frank M. Csultis, "Financial document processing system", published 2011-01-18, issued 2011-01-18</span><span class="Z3988" title="ctx_ver=Z39.88-2004&rft_val_fmt=info%3Aofi%2Ffmt%3Akev%3Amtx%3Apatent&rft.number=US7873576B2&rft.cc=US&rft.title=Financial+document+processing+system&rft.inventor=John+E.+Jones&rft.date=2011-01-18&rft.pubdate=2011-01-18"><span style="display: none;"> </span></span></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text"><cite id="CITEREFBridgwater" class="citation web cs1">Bridgwater, Adrian. <a rel="nofollow" class="external text" href="https://www.forbes.com/sites/adrianbridgwater/2020/03/09/appian-adds-google-cloud-intelligence-to-low-code-automation-mix/">"Appian Adds Google Cloud Intelligence To Low-Code Automation Mix"</a>. <i>Forbes</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2021-04-21</span></span>.</cite></span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text"><cite id="CITEREFAdamoAttivissimoDi_NisioSpadavecchia2015" class="citation journal cs1">Adamo, Francesco; Attivissimo, Filippo; Di Nisio, Attilio; Spadavecchia, Maurizio (February 2015). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://www.sciencedirect.com/science/article/pii/S0263224114005016">"An automatic document processing system for medical data extraction"</a></span>. <i>Measurement</i>. <b>61</b>: <span class="nowrap">88–</span>99. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/2015Meas...61...88A">2015Meas...61...88A</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.measurement.2014.10.032">10.1016/j.measurement.2014.10.032</a><span class="reference-accessdate">. Retrieved <span class="nowrap">31 January</span> 2021</span>.</cite></span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-13">^</a></b></span> <span class="reference-text"><cite id="CITEREFChangwanSeong-IlWon_Joon2020" class="citation journal cs1">Changwan, Kim; Seong-Il, Lee; Won Joon, Cho (September 2020). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://www.sciencedirect.com/science/article/abs/pii/S1877051720301994">"Volumetric assessment of extrusion in medial meniscus posterior root tears through semi-automatic segmentation on 3-tesla magnetic resonance images"</a></span>. <i>Orthopaedics & Traumatology: Surgery & Research</i>. <b>101</b> (5): <span class="nowrap">963–</span>968. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.rcot.2020.06.003">10.1016/j.rcot.2020.06.003</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a> <a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:225215597">225215597</a><span class="reference-accessdate">. Retrieved <span class="nowrap">31 January</span> 2021</span>.</cite></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text"><cite id="CITEREFDespotovićBartWilfried2015" class="citation journal cs1">Despotović, Ivana; Bart, Goossens; Wilfried, Philips (1 March 2015). <a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4402572">"MRI Segmentation of the Human Brain: Challenges, Methods, and Applications"</a>. <i>Computational Intelligence Techniques in Medicine</i>. <b>2015</b>: <span class="nowrap">963–</span>968. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://doi.org/10.1155%2F2015%2F450341">10.1155/2015/450341</a></span>. <a href="PMC_(identifier)" class="mw-redirect" title="PMC (identifier)">PMC</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4402572">4402572</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/25945121">25945121</a>.</cite></span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-15">^</a></b></span> <span class="reference-text"><cite id="CITEREFPutzuaCaocciDi_Rubertoa2014" class="citation journal cs1">Putzua, Lorenzo; Caocci, Giovanni; Di Rubertoa, Cecilia (November 2014). <a rel="nofollow" class="external text" href="https://www.sciencedirect.com/science/article/pii/S0933365714001031">"Leucocyte classification for leukaemia detection using image processing techniques"</a>. <i>Artificial Intelligence in Medicine</i>. <b>63</b> (3): <span class="nowrap">179–</span>191. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2Fj.artmed.2014.09.002">10.1016/j.artmed.2014.09.002</a>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://hdl.handle.net/11584%2F94592">11584/94592</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a> <a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/25241903">25241903</a>.</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text"><cite id="CITEREFEhrmannRomanelloClematideStröbel2020" class="citation conference cs1">Ehrmann, Maud; Romanello, Matteo; Clematide, Simon; Ströbel, Phillip; Barman, Raphaël (2020). <a rel="nofollow" class="external text" href="https://www.zora.uzh.ch/id/eprint/191270/">"Language Resources for Historical Newspapers: the Impresso Collection"</a>. <i>Proceedings of the 12th Language Resources and Evaluation Conference</i>. Marseille, France. pp. <span class="nowrap">958–</span>968.</cite></span>
</li>
<li id="cite_note-cini_archive_digitization-17"><span class="mw-cite-backlink">^ <a href="#cite_ref-cini_archive_digitization_17-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-cini_archive_digitization_17-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFSeguinCostinerdi_LenardoKaplan2018" class="citation conference cs1">Seguin, Benoit; Costiner, Lisandra; di Lenardo, Isabella; Kaplan, Frédéric (April 1, 2018). <a rel="nofollow" class="external text" href="https://www.ingentaconnect.com/content/ist/ac/2018/00002018/00000001/art00001">"New Techniques for the Digitization of Art Historical Photographic Archives - the Case of the Cini Foundation in Venice"</a>. <i>Archiving 2018 Final Program and Proceedings</i>. Society for Imaging Science and Technology. pp. <span class="nowrap">1–</span>5. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.2352%2Fissn.2168-3204.2018.1.0.2">10.2352/issn.2168-3204.2018.1.0.2</a>.</cite></span>
</li>
<li id="cite_note-18"><span class="mw-cite-backlink"><b><a href="#cite_ref-18">^</a></b></span> <span class="reference-text"><cite id="CITEREFAres_Oliveiradi_LenardoTourencKaplan2019" class="citation conference cs1">Ares Oliveira, Sofia; di Lenardo, Isabella; Tourenc, Bastien; Kaplan, Frédéric (11 July 2019). <a rel="nofollow" class="external text" href="https://infoscience.epfl.ch/record/268282"><i>A deep learning approach to Cadastral Computing</i></a>. Digital Humanities Conference. Utrecht, Netherlands.</cite></span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text"><cite id="CITEREFPetitpierre2020" class="citation thesis cs1">Petitpierre, Rémi (July 2020). <a rel="nofollow" class="external text" href="https://www.researchgate.net/publication/343017681"><i>Neural networks for semantic segmentation of historical city maps: Cross-cultural performance and the impact of figurative diversity</i></a> (MSc). <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/2101.12478">2101.12478</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.13140%2FRG.2.2.10973.64484">10.13140/RG.2.2.10973.64484</a>.</cite></span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text"><cite id="CITEREFFujisawaNakanoKurino1992" class="citation journal cs1">Fujisawa, H.; Nakano, Y.; Kurino, K. (July 1992). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://ieeexplore.ieee.org/document/156471">"Segmentation methods for character recognition: from segmentation to document structure analysis"</a></span>. <i>Proceedings of the IEEE</i>. <b>80</b> (7): <span class="nowrap">1079–</span>1092. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2F5.156471">10.1109/5.156471</a><span class="reference-accessdate">. Retrieved <span class="nowrap">3 February</span> 2021</span>.</cite></span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text">
<cite id="CITEREFTangLeeSuen1996" class="citation journal cs1">Tang, Yuan Y.; Lee, Seong-Whan; Suen, Ching Y. (1996). <span class="id-lock-subscription" title="Paid subscription required"><a rel="nofollow" class="external text" href="https://www.sciencedirect.com/science/article/abs/pii/S0031320396000441">"Automatic document processing: a survey"</a></span>. <i>Pattern Recognition</i>. <b>29</b> (12): <span class="nowrap">1931–</span>1952. <a href="Bibcode_(identifier)" class="mw-redirect" title="Bibcode (identifier)">Bibcode</a>:<a rel="nofollow" class="external text" href="https://ui.adsabs.harvard.edu/abs/1996PatRe..29.1931T">1996PatRe..29.1931T</a>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1016%2FS0031-3203%2896%2900044-1">10.1016/S0031-3203(96)00044-1</a><span class="reference-accessdate">. Retrieved <span class="nowrap">3 February</span> 2021</span>.</cite></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text"><cite id="CITEREFAres_OliveiraSeguinKaplan2018" class="citation conference cs1">Ares Oliveira, Sofia; Seguin, Benoit; Kaplan, Frederic (5–8 August 2018). <a rel="nofollow" class="external text" href="https://ieeexplore.ieee.org/document/8563218"><i>dhSegment: A Generic Deep-Learning Approach for Document Segmentation</i></a>. 2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR). Niagara Falls, NY, USA: IEEE. <a href="ArXiv_(identifier)" class="mw-redirect" title="ArXiv (identifier)">arXiv</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://arxiv.org/abs/1804.10371">1804.10371</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FICFHR-2018.2018.00011">10.1109/ICFHR-2018.2018.00011</a>.</cite></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://artmyn.com/">"Revolutionary Scanning Technology for Art"</a>. <i>Artmyn</i><span class="reference-accessdate">. Retrieved <span class="nowrap">3 February</span> 2021</span>.</cite></span>
</li>
</ol></div></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-06-23" href="https://en.wikipedia.org/wiki/?title=Document_processing&oldid=1297043329">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>